HTML API: Add custom text decoder - #6387
Closed
dmsnell wants to merge 23 commits into
Closed
Conversation
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Trac ticket: Core-61072
Token Map Trac ticket: Core-60698
From #5337 takes the HTML text decoder.
Replaces WordPress/gutenberg#47040
Status
The code should be working now with this, and fully spec-compliant.
Tests are covered generally by the html5lib test suite.
Performance
After some initial testing this appears to be around 20% slower in its current state at decoding text values compared to using
html_entity_decode(). I tested against a set of 296,046 web pages at the root domain for a list of the top-ranked domains that I found online.The impact is quite marginal, adding around 60 µs per page. For the set of close to 300k pages that took the total runtime from 87s to 105s. I tested with the following main loop, using
microtime( true )before and after the loop to add to the total time in an attempt to eliminate the I/O wait time from the results. This is a worst-case scenario where decode every attribute and every text node. Again, in practice, WordPress would only likely experience a fraction of that 60 µs because it's not decodingevery text node and every attribute of the HTML it ships to a browser.I attempted to avoid string allocations and this raised a challenge:
strpos()doesn't provide a way to stop at a given index. This led me to try replacing it with a simple look to advance character by character until finding a&. This slowed it down to about 25% slower thanhtml_entity_decode()so I removed that and instead relied on usingstrpos()with the possibility that it scans much further past the end of the value. On the test set of data it was still faster.For comparison, I built a version that skips the
WP_Token_Mapand instead relies on a basic associative array whose keys are the character reference names and whose values are the replacements. This was 840% slower thanhtml_decode_entities()and increased the average page processing time by 2.175 ms. The token map is thus approximately 36x faster than the naive implementation.Pre-decoding
In an attempt to rely more on
html_entity_decode()I added a pre-decoding step that would handle all well-formed numeric character encodings. The logic here is that if we can use a quickpreg_replace_callback()pass to get as much into C-code as we can, by means ofhtml_entity_decode(), then maybe it would be worth it even with the additional pass.Unfortunately the results were instantly slower, adding another 20% slowdown in my first 100k domains under test. That is, it's over 40% slower than a pure
html_entity_decode()whereas the code without the pre-encoding step is only 20% slower.The Pre-Decoder
Faster integer decoding.
I attempted to parse the code point inline while scanning the digits in hopes to save some time computing, but this dramatically slowed down the interpret. I think that the per-character parsing is much slower than
intval().Faster digit detection.
I attempted to replace
strspn( $text, $numeric_digits )with a custom look examining each character for whether it was in the digit ranges, but this was just as slow as the custom integer decoder.Quick table lookup of group/small token indexing.
On the idea that looking up the group or small word in the lookup strings might be slow, given that it's required to iterate every time, I tried adding a patch to introduce an index table for direct lookup into where words of the given starting letter start, and if they even exist in the table at all.
Table-lookup patch
This did not introduce a measurable speedup or slowdown on the dataset of 300k HTML pages. While I believe that the table lookup could speed up certain workloads that are heavy with named character references, it does not justify itself on realistic data and so I'm leaving the patch out.
Metrics on character references.
From the same set of 296k webpages I counted the frequency of each character reference. This includes the full syntax, so if we were to have come across
9it would appear in the list. The linked file contains ANSII terminal codes, so view it throughcatorless -R.all-text-and-ref-counts.txt
Based on this data I added a special-case for
", , and&before calling into theWP_Token_Mapbut it didn't have a measurable impact on performance. I'm led to conclude from this that it's not those common character references slowing things down. Possibly it's the numeric character references.In another experiment I replaced my custom
code_point_to_utf8_bytes()function with a call tomb_chr(), and again the impact wasn't significant. That method performs the same computation within PHP that this application-level does, so this is not surprising.For clearer performance direction it's probably most helpful to profile a run of decoding and see where the CPU is spending its time. It appears to be fairly quick as it is in this patch.
Attempted alternatives
WP_Token_Mapimplementation.",&, and , as they might account for up to 70-80% of all named character references in practice. This didn't impact the runtime. Runtime is likely dominated by numeric character reference decoding.ifchecks, rearranging code for frequency-analysis of code points, and replaced the Windows-1252 remapping with direct replacement. In a test 10 million randomly generated numeric character references, this performed around 3-5% faster than in the branch, but in real tests I could not measure any impact. The micro-optimizations are likely inert in a real context.In my benchmark of decoding 10 million randomly-generated numeric character references about half the time is spent exclusively inside
read_character_reference()and the other half is spent incode_point_to_utf8_bytes().substr() + intval()with an unrolled table-lookup custom string-to-integer decoder. While that decoder performed significantly better than a native pure-PHP decoder, it was still noticeably slower thanintval().I'm led to believe that his is nearly optimal for a pure PHP solution.
Character-set detections.
The following CSV file is the result of surveying the
/path of popular domains. It includes detections of whether the given HTML found at that path is valid UTF8, valid Windows-1252, valid ASCII, and whether it's valid in its self-reported character sets.A 1 indicates that the HTML passes
mb_check_encoding()for the encoding of the given column. A 0 indicates that it doesn't. A missing value indicates that the site did not self-report to contain that encoding.Note that a site might self-report being encoded in multiple simultaneous and mutually-exclusive encodings.
charset-detections.csv
html5libtestsTests: 609, Assertions: 172, Failures: 63, Skipped: 435.Tests: 607, Assertions: 172, Skipped: 435.Tests that are now possible to run that previously weren't.
Differences from
html_entity_decode()PHP misses 720 character references
Æ & & Á Â À ⁡ Å ≔ Ã Ä ∖ ⌆ ℬ ≎ © © ℭ Ç ⊖ ∲ ” ’ ∯ ℂ ∳ ⅅ ∇ ˙ ` ⋄ ¨ ≐ ⇓ ⇔ ⫤ ⟸ ⟺ ⇒ ⇕ ∥ ↓ ↽ ⇁ Ð É Ê È ∈ ≂ ⇌ ℰ Ë ⅇ ▪ ∀ ℱ > > ≥ ⋛ ≧ ≷ ⩾ ≫ ℋ ≎ Í Î Ì ℑ ⋂ ⁣ ℐ Ï < < ℒ ⟨ ← ⇆ ⌈ ⇃ ↔ ⊣ ⊲ ↿ ↼ ⇐ ⇔ ⋚ ≦ ≶ ⩽ ⇚ ⟵ ⟸ ⟺ ⟹ ↙ ℒ ≪ ℳ ​ ​ ​ ​ ≫ ≪   ℕ ∦ ∉ ≂̸ ∄ ≯ ≱ ⩾̸ ≵ ≎̸ ≏̸ ⋪ ⋬ ≸ ≪̸ ⩽̸ ≴ ⊀ ∌ ⋫ ⊂⃒ ⊃⃒ ≄ ≇ ≉ ∤ Ñ Ó Ô Ò Ø Õ Ö ‾ ∂ ± ℌ ℙ ≺ ⪯ ≾ ∏ ∷ ∝ " " ℚ ⤐ ® ® ↠ ℜ ⇋ → ⇄ ⊢ ↦ ⊳ ⇀ ⇒ ⇛ ℛ ↱ ↓ ← → ↑ ∘ ⊓ ⊏ ⊐ ⊔ ⋐ ≻ ≽ ∋ ∑ ⋑ ⊃ ⊇ Þ ™ ∴ ∼ ≃ ≈ Ú Û Ù _ ⎵ ⋃ ↑ ⇅ ⥮ ⊥ ⇑ ↖ ϒ Ü ⋁ ‖ ∣ | ≀   ⋀ Ý á â ´ ´ æ à ℵ & ∠ Å ≈ ≊ å ≈ ≍ ã ä ≌ ∽ ⌅ ∵ ∵ ϶ ℬ ⨀ ★ ⋁ ⋀ ⧫ ▪ ⊥ ⊥ ─ ⊠ ‵ ˘ ¦ ⋍ • ≏ ≏ ˇ ç ¸ ¸ ¢ · ✓ ↺ ↻ ® Ⓢ ⊛ ⊚ ⊝ ≗ ♣ ≔ ∁ ≅ ∮ ∐ © ⋞ ↶ ⋟ ¤ ↷ ⋎ ⋏ ⇓ ‐ ˝ ⅆ ‡ ⇊ ° ⇂ ⋄ ♦ ¨ ϝ ÷ ÷ ⋇ ⌞ ≐ ∸ ∔ ⊡ ↓ ⇃ ⇂ ▿ ▾ ⇵ ⥯ ⩷ ≑ é ê ≕ ⅇ ≒ è ∅ ∅ ε ϵ ≖ ≂ ⪖ ⪕ ≡ ≓ ð ë ∃ ⋔ ½ ½ ¼ ¾ ≧ ⋛ ≥ ⩾ ⋙ ≩ ⪊ ⪈ ≳ > ⋗ ⪆ ⪌ ≷ ≳ ≩︀ ℋ ℏ ♥ ⤥ ⤦ ↩ ↪ ℏ í î ¡ ⇔ ì ⅈ ∭ ℑ ℑ ı ∈ ∫ ℤ ⊺ ⨼ ¿ ∈ ⁢ ï ϰ ⇐ ⪋ ⟨ « ⇤ { “ „ ≤ ← ↢ ⇇ ↔ ⇆ ↭ ⋋ ⋚ ≦ ⩽ ⪅ ≲ ⌊ ≶ ↽ ↼ ⎰ ≨ ⪉ ⪇ ⟦ ⟷ ⟼ ⟶ ↫ ◊ ⌟ ⇋ ↰ ≲ [ ‘ ‚ < ⋖ ⊴ ◂ ≨︀ ¯ ✠ ↦ ↧ ↤ ↥ ∡ µ * · · ⊟ … ∓ ∓ ∾ ⊸ ≫̸ ⇎ ⇏ ≉ ♮   ≠ ↗ ↗ ≢ ⤨ ∄ ≧̸ ≱ ≧̸ ⩾̸ ≯ ↮ ∋ ∋ ⇍ ↚ ≰ ≰ ≦̸ ⩽̸ ≮ ≮ ∤ ¬ ∉ ∌ ∦ ⋠ ⪯̸ ⊀ ⪯̸ ↛ ⋫ ⋭ ⊁ ⋡ ⪰̸ ∦ ≁ ≄ ∤ ∦ ⋢ ⋣ ⊈ ⊂⃒ ⊈ ⫅̸ ⊁ ⪰̸ ⫆̸ ⊉ ⊉ ≹ ñ ⋪ ⋬ ⋭ ↖ ó ô ⊙ ò Ω ∮ ⊕ ℴ ª º ℴ ø õ ⊗ ö ∥ ¶ ∥ ϕ ℳ ℏ ⊞ ± ± £ ≺ ⪷ ≼ ⪯ ≼ ⪵ ⪹ ⋨ ∝ ≾ ⨌ ℍ ≟ " ⇒ ⤏ √ ⟩ ⟩ » → ⇥ ↬ ⤍ } ] ⌉ ” ℜ ℜ ℝ ® ⌋ → ↣ ⇁ ⇀ ⇉ ↝ ⋌ ⇄ ⇌ ⎱ ⟧ ’ ⊵ ▸ ≻ ⪸ ≽ ⪰ ⪺ ≿ ↘ ↘ § ∖ ∖ ⌢ ∣ ­ ς ≃ ← ∖ ∣ ♠ ∥ ⊑ ⊏ ⊑ ⊐ ⊒ ⊒ □ □ ▪ ⌣ ⋆ ¯ ⊆ ⫋ ⊊ ⊂ ⊆ ⫅ ⪰ ⪶ ⋩ ≿ ¹ ² ³ ⫆ ⊋ ⊃ ⊇ ⫌ ↙ ß ⎴ ⃛ ∴ ϑ ≈ ∼   ≈ ∼ þ ˜ × ⊤ ⤩ ◃ ⊴ ▹ ⊵ ≜ ≬ ↞ ⇑ ú û ù ↾ ⌜ ¨ ¨ ↑ ↕ ↿ ↾ ⊎ υ ⌝ ▵ ▴ ⇈ ü ⇕ ⊨ ϵ ∅ ϕ ϖ ∝ ↕ ϱ ς ⊊︀ ⫋︀ ⊋︀ ϑ ⊳ ∨ | ⊲ ⊃⃒ ∝ ⫌︀ ∧ ℘ ≀ ⋂ ◯ ⋃ ▽ ⟷ ⟵ ⨁ ⨂ ⟹ ⟶ ⨆ ⨄ △ ý ¥ ÿ ℨIn this list are many named character references without a trailing
;. This is because HTML does not require one in all cases. There's another behavior concerning numeric character references where the trailing;isn't required at certain boundaries.Further, whether or not the trailing
;is required is subject to the ambiguous ampersand rule, which guards a legacy behavior for certain query args in URL attributes which weren't properly encoded.Outputs from this PR
The
FromPHPcolumn shows howhtml_entity_decode( $input, ENT_QUOTES | ENT_SUBSTITUTE | ENT_HTML5 )would decode the input.The
DataandAttributecolumns show how the HTML API decodes the text in the context of markup (data) and of an attribute value (attribute). These are different in HTML, and unfortunately PHP does not provide a way to differentiate them. The main difference is in the so-called "ambiguous ampersand" rule which allows many "entities" to be written without the terminating semicolon;(though not all of the named character references may do this). In attributes, however, some of these can look like URL query params. E.g. is¬=dogssupposed to be¬=dogsor a query arg namednotwhose value isdogs? HTML chose to ensure the safety of URLs and forbid decoding character references in these ambiguous cases.Outputs from a browser
I've compared Firefox and Safari. The middle column shows the data value and the right column has extracted the
titleattribute of the input and set it as theinnerHTMLof theTD.The empty boxes represent unrendered Unicode charactered. While some characters, like the null byte, are replaced with a Replacement Character
�, "non-characters" are passed through, even though they are parser errors.Trac ticket:
This Pull Request is for code review only. Please keep all other discussion in the Trac ticket. Do not merge this Pull Request. See GitHub Pull Requests for Code Review in the Core Handbook for more details.